Phase 4: Neural Networks Lesson 1 of 5

Neural Networks:
The Building Blocks

Everything in modern AI, from image recognition to the chatbots you use today, is built on one architecture: the neural network. This lesson explains exactly how it works from the ground up, one neuron at a time.

You will learn
What an artificial neuron actually computes
What activation functions do and why they matter
How layers of neurons form a network
How data flows forward through a network
Building your first network in Keras

Loosely inspired by the brain

The term "neural network" comes from neuroscience. In 1943, Warren McCulloch and Walter Pitts published a mathematical model of how biological neurons might compute. In 1958, Frank Rosenblatt built on this to create the perceptron, the first trainable artificial neuron. These are the historical roots of modern deep learning.

That said, an important clarification for anyone who wants to be accurate: artificial neural networks are only loosely inspired by biology. A real biological neuron is enormously complex, with thousands of chemical signals, temporal dynamics, and structural properties that no artificial model comes close to replicating. When researchers say "inspired by the brain," they mean the basic idea of connected units that signal each other, not a faithful simulation of neuroscience.

So set aside the biology. An artificial neuron is a mathematical function. Understand it as maths and everything will make sense.

"Deep learning is a class of machine learning algorithms that use multiple layers to progressively extract higher-level features from raw input."

LeCun, Bengio and Hinton, Nature 2015 — the paper that defined the field for a generation

What one artificial neuron actually does

A single artificial neuron takes a set of inputs, multiplies each one by a weight, adds them all together along with a constant called the bias, and then passes the result through an activation function. That is the complete description of a neuron's computation.

A single artificial neuron: inputs, weights, sum, and activation
x₁ x₂ x₃ w₁ w₂ w₃ +b Σ activation f(z) y z = w₁x₁ + w₂x₂ + w₃x₃ + b output y = f(z)

The neuron computes z (the weighted sum plus bias), then applies the activation function f to get the output. Without the activation function, you could collapse any number of layers into a single multiplication and the network would have no more power than a basic linear model.

The weights determine how much each input contributes. A large positive weight means "when this input is high, the output should be high." A negative weight means "when this input is high, the output should be low." The bias is an offset that lets the neuron fire even when all inputs are zero, making the model more flexible.

These weights and the bias are the parameters of the neuron. Before training they are random. After training, they encode the pattern the network learned.

Activation functions: introducing non-linearity

The activation function is what makes neural networks powerful. Without it, stacking multiple layers of neurons would be mathematically identical to having just one layer. No matter how deep your network, it could only learn linear relationships, which would make it no better than linear regression.

By applying a non-linear function after each summation, you allow the network to learn curved, complex, non-linear relationships between inputs and outputs. This is the key insight.

Sigmoid
f(z) = 1 / (1 + e⁻ᶻ)
Output range: (0, 1)
Squashes any input to a value between 0 and 1. Historically the default choice, but has largely been replaced in hidden layers because it causes the vanishing gradient problem in deep networks.
Used in: binary output layers
ReLU
f(z) = max(0, z)
Output range: [0, +∞)
Returns the input if positive, otherwise returns zero. Simple, fast, and very effective. The default choice for hidden layers in most modern networks since 2012. Avoids the vanishing gradient problem that plagues sigmoid in deep networks.
Used in: most hidden layers
Softmax
f(zᵢ) = eᶻⁱ / Σeᶻʲ
Output range: (0, 1), sums to 1
Converts a vector of numbers into a probability distribution. If you have 10 output classes, softmax ensures all 10 outputs are positive and sum to exactly 1.0, making them interpretable as probabilities.
Used in: multi-class output layers
Why ReLU became the standard

Before ReLU, networks used sigmoid and tanh activations throughout. These functions both "squash" large values, which means gradients (the signals used for learning) become extremely small as they flow back through many layers. The network stops learning effectively. ReLU avoids this because for positive inputs, the gradient is always exactly 1, so signals can flow freely through many layers. The widespread shift to ReLU around 2012 was a major reason why very deep networks became trainable for the first time.

Layers: organising neurons into a network

One neuron is not very useful on its own. The power comes from connecting many neurons together in layers, with the output of one layer feeding into the input of the next. This creates a feedforward neural network, also called a multilayer perceptron (MLP).

A feedforward neural network: input layer, two hidden layers, output layer
INPUT HIDDEN 1 HIDDEN 2 OUTPUT x₁ x₂ x₃ y₁ y₂ 3 neurons 4 neurons 4 neurons 2 neurons

Every neuron in each layer is connected to every neuron in the next layer (this is called a "fully connected" or "dense" layer). The input layer simply receives the data. The output layer produces the final prediction. The hidden layers do the work of learning intermediate representations.

The number of layers and the number of neurons per layer are both hyperparameters that you choose before training. A network with two or more hidden layers is commonly called a deep neural network, which is where the term "deep learning" comes from. Depth is not a precise technical threshold; it is a term that reflects the shift in the field toward architectures with many layers.

The forward pass: how data flows through the network

When you feed a data point into a trained neural network to get a prediction, the computation that happens is called the forward pass. It is called "forward" because information flows in one direction: from the input layer, through each hidden layer in sequence, to the output layer. Nothing goes backwards during inference.

1
Input enters the network
Your raw data (a row of numbers) is fed into the input layer. Each feature becomes the activation of one input neuron. No computation happens here; the input layer just passes the data on.
2
Hidden layer 1 computes
Each neuron in the first hidden layer computes z = (weighted sum of inputs) + bias, then applies its activation function to get an output. These outputs are passed to the next layer.
3
Hidden layer 2 (and more) compute
Each subsequent hidden layer repeats the same process, building on the outputs of the previous layer. Early layers tend to learn simple patterns; later layers combine these into more complex representations.
4
Output layer produces the prediction
The output layer applies a final activation: sigmoid for binary classification, softmax for multi-class classification, or no activation (linear) for regression. This gives you the model's prediction.
Analogy

Think of a large company processing a job application. The document first reaches the HR team who extract key facts (education, years of experience). Their summary goes to the hiring manager who assesses fit for the role. That assessment goes to the department head who makes a final recommendation. Each layer processes what the previous layer passed on, adding a higher level of interpretation. The neural network does the same thing with numbers.

What makes neural networks so powerful

In 1989, mathematician George Cybenko proved something remarkable: a neural network with just one hidden layer containing enough neurons can approximate any continuous mathematical function to any desired degree of accuracy. This result, known as the Universal Approximation Theorem, tells us that the architecture is not the limiting factor. A sufficiently large network can, in principle, learn any pattern that exists in data.

This does not mean "any network learns anything." The theorem tells us about theoretical capacity, not about whether training will actually find the right weights, or whether you have enough data, or whether the network will generalise. But it does explain why the architecture is so widely applicable: image recognition, language translation, game playing, weather forecasting, protein structure prediction. One architecture, tuned differently, does all of it.

Your first neural network in Keras

Keras is the standard high-level interface for building neural networks. It ships as part of TensorFlow and is the fastest way to go from idea to working model. The API maps directly onto the concepts you just learned: you stack layers, specify their size and activation, then compile and train.

Python first_neural_network.py
import numpy as np
from tensorflow import keras
from tensorflow.keras import layers
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler

# Load and prepare data
X, y = load_breast_cancer(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42, stratify=y
)

# Scale inputs: neural networks are sensitive to feature scale
scaler = StandardScaler()
X_train = scaler.fit_transform(X_train)
X_test  = scaler.transform(X_test)

# Build the network
model = keras.Sequential([
    layers.Dense(64, activation='relu', input_shape=(X_train.shape[1],)),
    layers.Dense(32, activation='relu'),
    layers.Dense(1,  activation='sigmoid')
])

# Compile: choose loss function and optimiser
model.compile(
    optimizer='adam',
    loss='binary_crossentropy',
    metrics=['accuracy']
)

# Train
history = model.fit(
    X_train, y_train,
    epochs=50,
    batch_size=32,
    validation_split=0.15,
    verbose=0
)

# Evaluate on the held-out test set
loss, accuracy = model.evaluate(X_test, y_test, verbose=0)
print(f"Test accuracy: {accuracy:.4f}")
Typical output
Test accuracy: 0.9737

Notice a few things in the code. The network has an input layer (defined by input_shape), two hidden layers with ReLU activation, and one output neuron with sigmoid (because this is binary classification: benign or malignant). The loss function is binary cross-entropy, the standard choice for binary classification. The optimiser is Adam, a modern gradient descent variant that adapts the learning rate automatically and works well as a default.

Also notice the StandardScaler applied to the inputs. Neural networks are sensitive to the scale of their inputs in a way that decision trees and random forests are not. Features with large numeric ranges can dominate the gradient updates and make training unstable. Scaling all features to have mean zero and standard deviation one is standard practice before feeding data into a neural network.

Always scale your inputs for neural networks

This is not optional. If your features have very different scales (for example, age in the range 18-80, and income in the range 20,000-500,000), the income feature will produce gradients hundreds of times larger than the age feature. The network will pay almost no attention to age during training. StandardScaler (or MinMaxScaler for inputs bounded between 0 and 1) fixes this.

Hands-on activity

Build, break, and understand your first network

You are going to build a neural network from scratch in Keras, then deliberately change one thing at a time to understand what each part does. The goal is intuition, not just a working model.

01 Open the Lesson 4.1 Colab notebook. Run the breast cancer classifier exactly as shown. Record the test accuracy.
02 Remove the StandardScaler (skip the scaling step entirely). Retrain. How much does accuracy drop? This shows you why scaling matters.
03 Change all 'relu' activations to 'sigmoid'. Retrain. Does performance change? Does training take longer?
04 Add a third hidden layer with 16 neurons between the existing two. Does a deeper network help? What if you train for 100 epochs instead of 50?
05 Call model.summary() after building the network. Count the total number of parameters. Trace where each number comes from: (inputs × neurons) + bias for each layer.
Your Notes
Studying independently? Write your thoughts or answers below. Notes save automatically to your browser.
Practice Notebook
Run this lesson's code live in Google Colab
All examples + challenge exercises · Free GPU included · No setup required
Open In Colab
Progress
Done with this lesson?
Mark it complete to track your progress.